Journal of Molecular Evolution
○ Springer Science and Business Media LLC
Preprints posted in the last 30 days, ranked by how well they match Journal of Molecular Evolution's content profile, based on 22 papers previously published here. The average preprint has a 0.02% match score for this journal, so anything above that is already an above-average fit.
Thon, F. M.; Wittmann, M. J.
Show abstract
1. Plants produce a great chemodiversity, which is the diversity of specialized metabolites (SMs). These SMs are produced in complex metabolic pathways and play an important role in inter-species interactions. There are numerous hypotheses about the evolutionary processes which brought about and maintain chemodiversity. Some have been partially tested in lab and field studies. However, some of their assumptions and predictions are better tested by quantitative modeling, and so far no quantitative model has investigated the role of metabolic pathways. 2. To close this gap, we developed an individual-based model for metabolic pathway evolution. It models enzymes creating metabolites with various modifications. Enzymes undergo inheritance and mutation. We used the model to compare the screening and interaction diversity hypotheses. 3. The screening hypothesis predicts promiscuous enzymes, genetic drift, the presence of many non-beneficial metabolites, and high metabolite richness. The interaction diversity hypothesis predicts specialized enzymes, selection, the almost exclusive presence of beneficial metabolites, and situation- dependent metabolite richness. We found that the patterns predicted by the screening hypothesis did not occur, while those predicted by the interaction diversity hypothesis did. 4. This provides reason to favor the interaction diversity hypothesis over the screening hypothesis when connecting empirical results to their evolutionary context
Zhang, Z.; Xu, Y.
Show abstract
This study aims to quantify the genetic similarity of different species (from fish to humans) to the human reference genome (pp6, Homo sapiens.GRCh38) based on the allele presence/absence patterns of 33 language/cognition related gene SNV loci, identify key breakpoints during evolution, and evaluate the enrichment of language and cognition genes at these breakpoints. We designed a similarity calculation method relying on binary features (four columns for A/T/C/G), adopted five difference/distance measures (Sorensen, Rogers, Nei, Reynolds, and Hellinger), and converted them into similarity values (1/(1+distance)). For each method, samples were independently ranked, the first derivative of similarity was computed, and the top 12 peaks were selected as candidate breakpoints. Results show that the similarity curves from the five methods are highly consistent (correlation coefficients >0.9), with major peaks concentrated at positions 355, 363, 381, 382, 390, 400, etc., where the corresponding samples are predominantly ancient hominins and primates. Furthermore, we defined 13 peak groups (starting positions 355-401). For each peak within a group, pairwise SNV differences between the peak apex sample and its immediate left neighbor were compared, and the intersection F_INTERSECTION (shared differential loci) was obtained. For each F_INTERSECTION, we calculated the proportions of language genes and cognition genes. In addition, we computed the differential sets between adjacent groups' F_INTERSECTION to trace the gradual emergence of new loci. In F_INTERSECTION, language genes accounted for an average of 59.5%, and cognition genes for an average of 62.9%. The proportion of language genes reached a peak at position 383 (61.2%), while cognition genes peaked at position 386 (64.9%). High frequency peak samples include c25, c27, and ja2, suggesting that language cognition genes may have undergone independent intensification during Eurasian evolution. Differential analysis between adjacent F_INTERSECTION revealed a stepwise acquisition of new loci from position 355 to 401, with three bursts of newly added loci along the entire evolutionary axis. This study provides a quantitative framework based on similarity curves, offers a novel molecular perspective for understanding the evolution of language and cognitive abilities, and highlights the potential importance of East Asian archaic hominins in the evolution of language cognition genes.
Nguyen Huy, T.; Dong, Y.; Ly-Trong, N.; Vinh, L. S.; Minh, B. Q.
Show abstract
Model selection is a fundamental step in phylogenetic analysis that determines the best-fit model of sequence evolution for a given multiple sequence alignment. Popular model selection methods, such as ModelFinder, rely on statistical information criteria, such as the Bayesian Information Criterion (BIC) or the Akaike Information Criterion (AIC). However, these approaches are computationally expensive and the use of information criteria has been the subject of ongoing discussion. Recently, machine learning has emerged as a promising approach for phylogenetic model selection in both nucleotide and protein sequence analyses. ModelDetector is currently the only machine learning-based method for amino acid substitution model selection. However, because ModelDetector was trained on simulated data, it does not perform well on real datasets. Another limitation is that it does not support different rate heterogeneity across sites (RHAS) models. To overcome these limitations, we introduce ProtFinder, an efficient machine learning framework for protein model selection that predicts amino acid substitution models, RHAS models, and amino acid frequency models. To enable ProtFinder to work with real datasets, we employed a transfer learning strategy consisting of three stages: (1) initial training on large-scale simulated data, (2) joint training on both simulated and real data, and (3) final fine-tuning using real data only. Experimental results show that ProtFinder outperformed ModelDetector in amino acid substitution model selection. ProtFinder achieved comparable accuracy to the maximum likelihood method ModelFinder for substitution model selection on medium and large MSAs. It performs slightly better than ModelFinder in RHAS model selection and substantially outperforms it in amino acid frequency model determination. Notably, ProtFinder is up to 1,400 times faster than ModelFinder in terms of inference time, making it particularly suitable for medium and large datasets.
Lawson, M. E.; Sanow, K.; Fratian, M.; Matura, M.; Scanlon, R.; Richard, M.; Nakhla, M.; Rele, C. P.; Thompson, J. S.; Findlay, G. D.; O'Rourke, K. S.
Show abstract
Gene model for the ortholog of Density regulated protein (DENR) in the Apr. 2013 (BCM-HGSC Dpse_3.0/DpseGB3) Genome Assembly (GenBank Accession: GCA_000001765.2) of Drosophila pseudoobscura. This ortholog was characterized as part of a developing dataset to study the evolution of the Insulin/insulin-like growth factor signaling pathway (IIS) across the genus Drosophila using the Genomics Education Partnership gene annotation protocol for Course-based Undergraduate Research Experiences.
Saha, A.; Ghosh, A.; Majumdar, S.
Show abstract
THAP9 is a transposable element-derived gene which encodes a protein that is homologous to the active Drosophila P-element transposase (DmTNP). Both THAP9 and DmTNP possess a C-terminal domain (CTD) which is functionally uncharacterized. Sequence and structural analysis suggest that the THAP9-CTD has a novel fold which is only found in THAP9 homologs. To explore the evolutionary history and characteristics of this novel domain, exhaustive phylogenetic analysis (using MSA, structure prediction, MSTA-based clustering) was performed. THAP9-CTD homologs were more widely distributed throughout the animal kingdom in comparison to DmTNP-CTD homologs which were restricted to arthropods. Moreover, the THAP9-CTD homologs were more conserved, especially among mammals and birds and their average length increased in a class-specific manner. Comparison with the DmTNP-CTD homologs demonstrates that although their respective CTDs may have evolved independently, they both surprisingly share similar secondary structure elements consisting of three conserved helical regions made of hydrophobic residues that are predicted to make up a conserved core. The role of the respective CTDs were further investigated by creating truncation mutants lacking the CTD. Interestingly both THAP9 and DmTNP truncation mutants are still capable of DNA excision and integration suggesting that their respective CTDs are not essential for DNA transposition. Moreover, CTD truncation favours DNA integration in THAP9: this suggests that CTD acquisition during evolution may have led to THAP9 domestication as observed in other transposable element-derived genes like Rag1 and piggybac, which have similar terminal regulatory domains.
Ahammed, K. S.; Miramon, P.; Schrettenbrunner, L.; Cruz, M. R.; Huh, E. Y.; Hu, H.; Israni, B.; Wilson, H. B.; Li, Z.; Lee, S. C.; Blango, M. G.; Garsin, D. A.; Lorenz, M. C.; van Hoof, A.
Show abstract
The majority of eukaryotes encode some intron-containing pre-tRNAs. Splicing of these pre-tRNAs requires a dedicated tRNA splicing machinery. The fungal and trypanosome tRNA ligase, Trl1, and the human RNA ligase, RTCB, catalyze an essential step in tRNA splicing. However, Trl1 and RTCB are nonhomologous and biochemically and structurally distinct from each other. Therefore, Trl1 could serve as a broad-spectrum antifungal and anti-trypanosomal target. While the functions and requirements of the three catalytic Trl1 domains have been extensively characterized in the model yeast Saccharomyces cerevisiae, the roles of Trl1 orthologs in pathogenic fungi remain unexplored. Here, we validate Trl1 as one of the few promising novel drug targets for the development of antifungal therapeutics. Functional analyses of the three Trl1 domains show that only the "sealing" domain is essential for growth and viability in Candida albicans and Aspergillus fumigatus. In contrast, the two "healing" domains are dispensable in these pathogenic fungi, suggesting the presence of redundant healing enzymes, unlike in S. cerevisiae. These findings indicate that only the sealing domain is a good drug target. Our analysis also shows that the Mucor enzyme, which only contains the sealing domain, is essential. Using a Caenorhabditis elegans infection model of C. albicans, we further demonstrated that inhibiting Trl1 expression protects worms during an established infection. In contrast to these fungal pathogens, we show that all three domains of Trl1 are essential in Trypanosoma brucei. Our findings show that the essentiality of the Trl1 sealing is conserved in important human pathogens and provides an impetus for future drug development. SIGNIFICANCEFungal infections are an important cause of human disease and death and difficult to treat and there is an urgent need to develop additional drugs. Based on studies in yeast, one promising target for antifungal drug development is the tRNA splicing pathway. Human tRNA ligase is fundamentally distinct from the fungal one. To investigate the possibility of developing tRNA ligase-targeting drugs, we investigated the function of the catalytic domains of fungal tRNA ligase in different fungal pathogens. Surprisingly, only the first domain is essential in these pathogens and yeast is not a good model fungus. In contrast, all three domains of Trypanosome tRNA ligase are essential. These findings provide an impetus for future drug development.
Lieser, B. C.; Laskowski, L. F.; Huber, R.; Kolker, K. O.; Arsham, A. M.; Rele, C. P.; Toering Peters, S.
Show abstract
Gene model for the ortholog of Insulin-like peptide 3 (Ilp3) in the D. pseudoobscura Apr. 2013 (BCM-HGSC Dpse_3.0/DpseGB3) Genome Assembly (GenBank Accession: GCA_000001765.2) of Drosophila pseudoobscura. This ortholog was characterized as part of a developing dataset to study the evolution of the Insulin/insulin-like growth factor signaling pathway (IIS) across the genus Drosophila using the Genomics Education Partnership gene annotation protocol for Course-based Undergraduate Research Experiences.
Balasov, M.; Shibata, E.; Akhmetova, K.; Dutta, A.; Chesnokov, I.
Show abstract
In eukaryotes, DNA replication requires the origin recognition complex (ORC), a six-subunit assembly that promotes replisome formation on chromosomal origins. Orc6 is the smallest and least evolutionarily conserved among all ORC subunits. In Drosophila, Orc6 binds tightly with the core ORC(1-5) and is required for DNA binding and replication initiation, whereas in Xenopus and human systems Orc6 loosely associates with the rest of the complex resulting in some differences for replication-associated activities. Despite these variations, Orc6 remains essential for viability in all species. In current study we analyzed specific residues within the C-terminal 11 helix that is critical for stable association of Orc6 with the ORC complex in Drosophila. Human Orc6 lacks these residues, however it possesses a strong nuclear localization signal (NLS) that is absent in Drosophilidae. We propose that this NLS drives human protein to the nucleus and compensates for weaker Orc6-ORC(1-5) interactions by increasing the nuclear concentration of Orc6 and shifting the equilibrium toward formation of the fully assembled ORC complex at the DNA.
Gorstein, E.; Tang, M.; Bruzzone, H.; Solis-Lemus, C.
Show abstract
Standard methods for ancestral sequence reconstruction (ASR) rely on substitution models for the residues in a biological sequence and assume independent evolution across these sites, ignoring the epistatic interactions that shape molecular evolution. In contrast, deep learning models like variational autoencoders (VAEs) can learn low-dimensional representations ("embeddings") of sequences in a protein family that may implicitly handle these dependencies, raising the possibility of performing more accurate ASR by interpolating between extant sequence embeddings within the VAE's latent space. In this study, we test this hypothesis by developing and evaluating a VAE-based ASR pipeline. Benchmarking this approach against established likelihood-based and parsimony methods using various simulations of protein evolution, including scenarios with and without epistasis, we find that the VAE-based approach is consistently and significantly outperformed by standard methods, even in epistatic regimes where it was hypothesized to have an advantage. We further show that this failure is not due to a lack of phylogenetic structure in the latent space, which does contain evolutionary signal. Rather, the primary limitation is the information loss inherent to the autoencoding process: the VAE's decoder cannot generate sequences with sufficient fidelity for the precise demands of ASR.
Samo, N.; Nguyen, L.; Kumawat, S.; Choi, J. Y.
Show abstract
Telomeres are nucleoprotein structures that protect chromosome ends and are maintained by the Telomerase Reverse Transcriptase (TERT) protein that uses a noncoding Telomerase RNA (TR) as a template. In monkeyflowers, Mimulus lewisii had an ancient TR gene duplication, synthesizing an evolutionarily atypical sequence heterogeneous telomere. How TERT interacts with both TR paralogs during telomere maintenance is unknown and answers can shed novel insights underlying telomere function. Using new genome assemblies we discovered TERT is rapidly evolving in lineages sharing the TR duplication. We investigated the functional consequences arising from the rapid evolution, first by using yeast three-hybrid and testing the physical binding between conspecific and heterospecific TERT-TR combinations. Results showed TERT binds both ancestral (TR1) and derived (TR2) TR paralogs in M. lewisii, but not in species without a functioning TR2. We located the region of TR binding to amino acids near the KRxR motif. We then combined next-generation sequencing with Telomeric Repeat Amplification Protocol and discovered M. lewisii had high telomerase activity. Comparative transcriptomics indicated no strong evidence of expression divergence in telomere maintenance genes for M. lewisii, suggesting rapid evolution shaped TERT protein sequence. In vivo activity of M. lewisii telomerase was investigated by analyzing F1 telomeres generated by crossing M. lewisii and M. verbenaceus, which doesnt have a functioning TR2. Results showed M. verbenaceus chromosome ends in the F1 had converted into M. lewisii telomeres, suggesting dominance of the M. lewisii telomerase. We demonstrate TERT-TR coevolution can have significant consequences on the evolution of plant telomeres. Significance statementTelomeres protect chromosome ends and are maintained by the telomerase complex. We discovered the catalytic component of the telomerase (TERT) was rapidly evolving in monkeyflowers (Mimulus) and studied the molecular consequences. In M. lewisii, TERT evolved lineage-specific amino acids to bind two sequence divergent telomerase RNA paralogs. Telomerase activity assay showed M. lewisii synthesized more telomere repeats compared to its sister species without the TR duplication, and transcriptomics indicated this was not due to a change in telomere maintenance gene expression. Genetic experiments in interspecies hybrids showed M. lewisii telomerase could convert chromosome ends in sister species into M. lewisii-like telomeres suggesting functional dominance. We show rapid evolution of the telomerase can have significant effects on telomere evolution.
Matsuda, T.; Yokogawa, T.; Hidetaka, S.; Sora, M.; Ihara, A.; Toba, A.; Kawai, K.; Norimoto, G.; Hirata, A.; Hori, H.; Yamagami, R.
Show abstract
N2-methylguanosine (m2G) is widely found at multiple positions in tRNAs across the three domains of life. Tryptophan tRNA from Thermococcus kodakarensis contains m2G at position 67. We previously proposed that the tRNA m2G methyltransferase Trm14 is responsible for m2G67 formation in tRNATrp from T. kodakarensis, although Trm14 was originally identified as the enzyme catalyzing m2G6 formation in tRNACys in Methanocaldococcus jannaschii. Thus, it remained unclear whether Trm14 could also methylate G67. Here, we characterized archaeal Trm14. Biochemical analyses using recombinant T. kodakarensis Trm14 revealed that the enzyme catalyzes m2G formation at positions 6 and 67 in T. kodakarensis tRNACys and tRNATrp transcripts, respectively. Mass spectrometric analyses demonstrated the loss of m2G6 and m2G67 in native tRNACys and tRNATrp, respectively, from a T. kodakarensis trm14 gene disruptant strain, providing direct evidence for the dual-site specificity of T. kodakarensis Trm14. The growth phenotype of the trm14 gene disruptant strain was comparable to that of the wild-type strain. In contrast, a trm14/trm11 double disruptant, in which trm11 encodes the tRNA m2G10/m22G10 methyltransferase, exhibited severe growth retardation at 95 {degrees}C. This suggests that m2G6/m2G67 and m2G10/m22G10 cooperatively contribute to cellular fitness at high temperatures. Biochemical analyses revealed that Trm14 methylates all 46 T. kodakarensis tRNA transcripts. Furthermore, we found that recombinant M. jannaschii Trm14 methylated both positions. In contrast, the bacterial ortholog TrmN modified only position 6 in tRNA. Overall, this study expands our understanding of archaeal Trm14 by demonstrating its broader substrate specificity and the physiological significance of these modifications under hyperthermophilic conditions.
Mathur, C.; Davis, E. T.; Ehrbar, D.; Omeoga, H. C.; Endres, L.; Byrne, S. R.; Begley, U.; Dedon, P. C.; Begley, T. J.
Show abstract
Oncogenes and tumor-suppressor genes play opposing roles in cancer biology to promote and restrict growth, respectively. Codon usage patterns interface with tRNA modifications to control translation, leading to gene-specific codon signatures with regulatory potential. As such, codon-biased translational regulation has been identified as a driver of proliferation and drug resistance in multiple cancers. We used advanced codon analytics methods to characterize and compare codon usage bias in oncogenes and tumor suppressor genes (TSGs) from humans and mice at group and gene-specific levels. We demonstrate that human oncogenes exhibit a distinct and opposing codon usage pattern to TSGs. This phenomenon is also present in mice but with less distinct oncogene bias relative to humans. Further comparison to 447 gene ontology groups demonstrated that human oncogenes have the most distinct codon usage patterns in the genome, while also highlighting that codon bias can separate functionally related genes and pathways from other biological processes. Using gene-specific codon analytics, we determined that human oncogenes have two types of extreme codon bias: a large group (N = 43) over-using G/C ending (GC3) codons and a smaller group (N = 12) over-using A/U (AU3) ending codons. While GC3 bias has been linked to increased translation in general, the AU3 finding suggests that genetic, environmental, or stress-related signals could drive the translation of this small group of oncogenes. The less extreme bias observed in mouse oncogenes and tumor suppressors likely underscores species-specific differences in oncogenic translation programs. Together, our findings highlight codon usage bias as a potential determinant of oncogene expression, provide a framework for ontology-based codon analysis, and uncover on species-specific differences in oncogene translation and codon usage biases.
McDonald, J. M. C.; Guo, Q.; Delgado, S.; Amendola, C. A.; Garg, I. A.; Reed, R. D.
Show abstract
Butterfly wings present a tremendous gallery of colorful patterns, offering a unique opportunity to study how developmental pattern formation processes evolve. We still do not understand the genetic basis of several key aspects of wing pattern development, however. Three paralogous POU domain transcription factors nubbin, ventral veinless (vvl), and pdm3 are all known wing development genes in Drosophila melanogaster. Here we combine gene expression and knockout approaches to show that each of these genes plays multiple novel wing patterning roles in the common buckeye butterfly, Junonia coenia. We found that nubbin controls eyespot pattern determination via a non-cell autonomous repressor-like effect originating at the wing veins, such that nubbin knockouts have larger eyespots. nubbin also regulates pigment identity and scale morphology across the wings. We also found that vvl regulates pigment identity of the discal bands and ventral hindwing. Last, we found that pdm3 is required for determining the outer rings of eyespot patterns, where it is co-expressed with spalt and the lncRNA ivory. pdm3 is also necessary for determining wing margin stripes, where it is again co-expressed with spalt, leading us to propose that the eyespot and wing margin gene regulatory networks could be homologous. Finally, pdm3 affects pigmentation of the ventral hindwing, phenocopying the seasonally-plastic color switch in J. coenia. Together, our work shows that POU domain transcription factors play diverse roles in butterfly wing pattern development and highlights nubbin as one of the first genes implicated in the repressive function of wing veins in color pattern determination. Highlights- Gene expression and knockouts reveal three POU factors regulate butterfly wing color pattern - nubbin regulates eyespot development, likely via a repressor from the wing veins - nubbin controls scale color and morphology across the wing - pdm3 coordinates eyespot development and is co-expressed with spalt and ivory - Expression of genes in the eyespot and wing margin suggests network homology
Lin, R.; Wang, C.; Hu, Y.; Wang, C.
Show abstract
The canonical genetic code is used by most known forms of life, yet explaining its historical origin and present day functional performance requires comparison with the enormous space of possible codon-to-output assignments. Here, the code is formulated as a hierarchy of constrained mapping problems spanning codeword length, degeneracy composition, synonymous-block partitioning, semantic assignment and a coarse decoder layer. Exact structural analyses identify triplets as a Pareto choice under a fixed-length full-codebook model and show that anonymous degeneracy statistics alone do not explain the canonical profile. Within a fixed canonical block architecture and under specified objective functions, recurrent AAindex-based rule learning contracts the 20! amino-acid assignment space to 2.72 ** 1011 admissible mappings, from which 108 complete codes are sampled. In this screened conditional candidate library, the standard genetic code ranks in the best 0.9749% under the equal-weight three-objective score and in the best 1.801% when accessible replacement diversity is added. Sensitivity analyses show that this position is broad across many, but not all, tested objective weights and aggregation rules. These results describe a conditional multi-objective compromise; they do not establish global optimality, historical inevitability or cellular feasibility of decoder redesign.
Varshney, D.; Tajjar, M. H.; de Vries, J.; Hutter, F.; Rensing, S. A.
Show abstract
How morphological complexity evolves is still enigmatic. While there is evidence in algae and plants as well as animals that diversification of the repertoire of transcription factors (TF) is causative for evolution of organismal complexity, there are many examples from lineages that follow their own way of complexity evolution, for example by expansion of particular families. For land plants, correlation of the size of the TF complement with number of cell types (as a proxy for morphological complexity) has been shown, and several families were identified as candidates to drive complexity evolution. Here, we expand a previously available dataset of cell type numbers from 12 to 82 proteomes and introduce a four class body plan scheme. We find that the total TF complement correlates with the number of cell types of Archaeplastida (primary plastid bearing plants and algae). We used TabPFN (Tabular Prior-data Fitted Network) for binary (uni- vs. multicellularity) as well as for four class Bauplan classification. TabPFN is able to predict the morphological complexity with high accuracy. This approach allows to determine organismal complexity based on the gene space of an organism. Based on our results, we can confirm that plant morphological evolution is driven by gain and expansion of TF families.
Beaulieu, J. M.; O'Meara, B. C.
Show abstract
Fossilized birth-death (FBD) models provide a powerful framework for estimating diversification from phylogenies that include fossil taxa. However, the original formulation makes a key assumption that sampled ancestors (k-type fossils) should be commonly observed. Beaulieu & OMeara (2023) showed that this assumption is often violated in empirical datasets, where fossils are represented mostly or entirely as extinct terminal taxa (m-type fossils), which can lead to biased parameter estimates. Here, we derive the fossilized birth-death of terminal fossils (FBDT) model, an extension of the FBD that accommodates incomplete fossil samples in which only terminal fossil occurrences are observed. We implement the model within the state-dependent speciation and extinction (SSE) framework and evaluate its performance using simulations spanning homogeneous and heterogeneous diversification scenarios. Across a wide range of fossil sampling rates, the FBDT model recovered diversification parameters that closely matched those obtained from complete fossil samples while avoiding the systematic biases that arise when sampled ancestors are unobserved. These results demonstrate that modifying the likelihood to reflect how fossil datasets are assembled provides a simple and effective extension of the FBD framework for many empirical applications.
Charles, B.; Moehring, A. J.
Show abstract
Mutant tRNA mistranslation is a phenomena in which specific tRNA gene mutations cause translational errors wherein amino acids incorporated during translation differ from those coded by mRNA sequences. The biological consequences of mutant tRNA mistranslation are highly complex and contextual, but are often deleterious. Importantly, many variables exist that are likely to affect the outcomes of any one mutant mistranslating tRNA; biochemical characteristics of the exchanged amino acids, frequency of use, identity of mistranslated products, and so on. Here, we generate a carefully curated array of tRNA mutants to assess which characteristics of mistranslating tRNA variants are most predictive of deleterious phenotypes. We find that toxicity often arises from mutant tRNAs that induce dramatic changes in biochemical properties between exchanged amino acids, as well as mistranslation that occurs more frequently. Importantly, exceptions are also observed to each of these general rules, implying instead that some deliriousness may arise from more granular, product-specific mechanisms. Additionally, quantification of mistranslation via mass spectrometry demonstrates a weak relationship between quantity of mistranslated products and severity of toxic phenotypes. Lastly, despite previous establishment of mutant mistranslating tRNA models in Drosophila melanogaster, incorporation of higher frequency mistranslating tRNA variants was largely unsuccessful, and incorporation of lower frequency mistranslating tRNA variants produces no detectable developmental phenotypes. These experiments support the notion that although general rules may be capable of reasonably predicting their consequences, each mistranslating tRNA variant warrants individual consideration and investigation for thorough understanding of its biological outcomes and precise mechanisms of toxicity.
DeMontigny, W. C.; Delwiche, C. F.
Show abstract
Selective pressures can vary across both sites and evolutionary lineages; however, most codon models accommodate heterogeneity along only one of these dimensions and require the number of selective regimes to be specified in advance. Here, we introduce OmegaSwitch, a Bayesian phylogenetic software framework for inferring changes in the nonsynonymous-to-synonymous substitution-rate ratio (dN/dS) across sites and through evolutionary time. We implement a Markov-modulated codon model in which lineages transition among discrete dN/dS regimes and use reversible-jump Markov chain Monte Carlo to infer the number of regimes simultaneously. We further develop a Dirichlet-process mixture extension that allows the parameters governing these time-heterogeneous processes to vary among sites. Ancestral sampling produces joint posterior distributions of dN/dS across sites and nodes of the phylogeny, enabling lineage- and site-specific summaries with quantified uncertainty. Simulation analyses showed that both the posterior intervals for dN/dS and the number of evolutionary regimes were well calibrated under both models. We demonstrate OmegaSwitch using vertebrate alpha- and beta-globins. OmegaSwitch therefore provides a flexible Bayesian framework for investigating how selective pressures vary across protein-coding sequences and phylogenetic history.
Tong, Y.; Rossetto Marcelino, V.; Turnbull, R. B.; Verbruggen, H.
Show abstract
Chloroplast or plastid genomes are essential resources for studying the evolution and diversity of algae and land plants. Although thousands of plastid genomes have been sequenced, their full potential has not been realised; derived resources such as orthogroup databases and reference datasets for metagenomic profiling remain underdeveloped. We present the ChlORIS database to address these problems across all algal phyla. From 2,254 publicly available algal plastid genomes, after dereplication we clustered 2,531 orthogroups from the annotated proteins and selected 496 orthogroups with consistent gene naming, enabling cross-genome comparisons of homologous plastid proteins. We further selected 224 core orthogroups, each containing more than 10 protein sequences, for which we produced score-calibrated hidden Markov models (HMMs), multiple sequence alignments and predicted protein structures. The value of these resources for phylogenomics is demonstrated through a large-scale plastid phylogeny of 859 taxa spanning all major algal lineages. We characterised the protein HMMs by cross-referencing them to Pfam domains and calibrated score cutoffs for reliable detection. The metagenomic database, HMM library, nucleotide and amino acid alignments, predicted structures and protein metadata, cross-linked to UniProt and InterPro (Pfam), are openly available on the ChlORIS website at https://chloris.codeberg.page/.
Vermette, O.; Mixoy, R. L.; Flynn, J. M.
Show abstract
Satellite DNA is long arrays of tandem repetitive DNA located often near the centromeres of chromosomes, whose function, or lack of, has been debated since its discovery. Although situated in heterochromatin, satellite DNA may be expressed as long noncoding RNAs (lncRNAs). Although there are a few examples of satellite lncRNAs being characterized, and functions suggested, how widespread and functionally important they may be for developmental processes is not understood. Here, we take an evolutionary approach to investigate satellite lncRNA expression in Drosophila spp. ovaries, a tissue whose development is well-characterized but where satellite expression has only been minimally explored. Using a publicly-available total RNAseq dataset, we find that 118/156 surveyed satellite DNAs were expressed across 10 species, with 33 satellites having high expression over 20 RPM. However, all but two of these expressed satellites (AAACTAC in D. virilis and ACAGACAGACAGG in D. ananassae) had higher read counts in a sister smallRNA dataset, suggesting that most satellite transcripts primarily serve as precursors for piRNA biogenesis. The two "stand-alone" lncRNAs were highly strand-biased, with 96-97% of the total reads coming from one strand. We further investigated AAACTAC expression with RNA FISH and found the transcript is specifically present in the oocyte nucleus following a dynamic spatiotemporal pattern, with the highest expression in stage 3-5 oocytes. The transcription pattern of AAACTAC is conserved in the three other virilis clade species that contain this satellite DNA. Further, we found expression of unrelated satellites in more distantly related D. borealis and littoralis both in the oocyte and the nurse cells. Overall, our work identifies a novel lncRNA AAACUAC found in the early oocyte nucleus, which is conserved across ~5 MY of evolution, and is therefore a strong candidate for the discovery of novel functions of satellite lncRNAs in development.